Characters, Glyphs and Beyond

نویسندگان

  • Tereza Haralambous
  • Yannis Haralambous
چکیده

The distinction between characters and glyphs is a fundamental issue of computing. This talk aims in giving a new definition of these notions. We first review and comment the definitions given in various standards. Then we give and explain our own definitions. We consider that the Unicode character model is lacunary and formulate a proposal for adding supplementary information and obtaining thus “rich Unicode characters.” We illustrate our arguments with many examples, taken from various writing systems. The distinction between characters and glyphs is currently a very popular issue. The complexity of this issue is, in some sense, related to the fact that computer systems have been build by engineers not very proficient in linguistics, and interested only in the English language. Exploring non-latin writing systems one realizes what has not been clear from the beginning: that modelizing written language is not a trivial task, and that it is fundamental to all exchange and processing of textual information. Let us start the exploration of this universe by giving some definitions of the terms we are using. Let us see how the terms “character” and “glyph” are defined. According to ISO 9541 [6] released in 1991, a “glyph” is “a recognizable abstract graphic symbol which is independent of any specific design,” while a “glyph image” is “an image of a glyph, as obtained from a glyph representation diplayed on a presentation surface,” where “glyph representation” is “the glyph shape and glyph metrics associated with a specific glyph in a font resource.” We may argue if this distinction between “abstract glyph” and “concrete glyph” is necessary, but this is how ISO 9541 defines these. According to W3C (quoting “A Character Model for the World Wide Web” by Martin Drst and others [2]), a character is “the smallest component of written language that has semantic values; refers to the abstract meaning and/or shape.” We find this definition quite vague since everything we perceive may or may not have semantic value, depending on our culture, context and even mood. . .We all know that Unicode is full of inconsistencies, because of its requirement to be compatible with legacy encodings. Has this definition been made to encompass Unicode

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Providing some UTF-8 support via inputenc

3 Mapping characters — based on font (glyph) encodings 11 3.1 About the table itself . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.2 The mapping table . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 3.3 Notes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.4 Mappings for OT1 glyphs . . . . . . . . . . . . . . . . . . . . . . . 24 3.5 Mappings for OMS g...

متن کامل

Learning Chinese Word Representations From Glyphs Of Characters

In this paper, we propose new methods to learn Chinese word representations. Chinese characters are composed of graphical components, which carry rich semantics. It is common for a Chinese learner to comprehend the meaning of a word from these graphical components. As a result, we propose models that enhance word representations by character glyphs. The character glyph features are directly lea...

متن کامل

Generation of Glyphs for Conveying Complex Information, with Application to Protein Representations

We present a method to generate glyphs which convey complex information in graphical form. A glyph has a linear geometry which is specified using geometric operations, each represented by characters nested in a string. This format allows several glyph strings to be concatenated, resulting in more complex geometries. We explore automatic generation of a large number of glyphs using a genetic alg...

متن کامل

Omega Becomes a Texteme Processor

The distinction between “characters” and “glyphs” is a rather new issue in computing, although the problem is as old as humanity: our species turns out to be a writing one because, amongst other things, our brain is able to interpret images as symbols belonging to a given writing system. Computers deal with text in a more abstract way. When we agree that, in computing, all possible “capital A” ...

متن کامل

ΩTimes and ΩHelvetica Fonts Under Development: Step One

ΩTimes and ΩHelvetica will be public domain virtual Timesand Helvetica-like fonts based upon real PostScript fonts, which we call “Glyph Containers”. They will contain all necessary characters for typesetting efficiently (that is, with TEX quality) in all languages and systems using the Latin, Greek, Cyrillic, Arabic, Hebrew and Tifinagh alphabets and their derivatives. All Unicode characters w...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2004